Papers with normalized version
Benefits of Data Augmentation for NMT-based Text Normalization of User-Generated Content (D19-55)
Copied to clipboard
| Challenge: | Social media texts are considered important language resources for several NLP tasks, but their use of non-standard words makes it difficult to process and analyze UGC. |
| Approach: | They propose to use a Neural Machine Translation approach to normalize lexical variants to their canonical forms to overcome performance drop in UGC. |
| Outcome: | The proposed approach overcomes a data bottleneck in Dutch, a low-resource language. |
Ranking Human and LLM Texts Using Locality Statistics (2026.findings-eacl)
Copied to clipboard
| Challenge: | The paper extends the Data Movement Distance (DMD) metric defined to measure the locality in computer memory to text by defining a new term designed to better characterize low-frequency tokens. |
| Approach: | They propose to define a normalized version of the Data Movement Distance (nDMD) term is designed to better characterize low-frequency tokens. |
| Outcome: | The proposed normalized version outperforms baselines and improves performance on the English subset of the M4 dataset and the GenAI detection shared task. |